104 results found.
Speech
Corpus,
Language Type:
Multilingual
Languages:
Amharic Bosnian Croatian Dari English French Georgian Haitian Hausa Hindi Korean Mandarin Chinese Persian Portuguese Pushto Russian Spanish Turkish Ukrainian Urdu Vietnamese Yue Chinese
Availability:
From Owner
License:
LDC
Size:
215 hoursProduction Status:
Existing-used
Use:
Language Identification
-
Paper title:Modeling and training strategies for language recognition systems
-
Paper track:4.1 Language identification and verification, lang/Oral Presentation
-
Paper status:Accept
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Raphaël Duroselle | 2009 NIST Language Recognition Evaluation Test Set | /N |
Documentation:
None
Speech
Corpus,
Language Type:
Multilingual
Languages:
Arabic Bengali Dari Egyptian Arabic English Georgian Hindi Iranian Persian Italian Japanese Khmer Korean Lao Mandarin Chinese Min Nan Chinese Moroccan Arabic Panjabi Persian Russian Spanish Tagalog Thai Tigrinya Urdu
Availability:
From Owner
License:
LDC
Size:
640 hoursProduction Status:
Existing-used
Use:
Language Identification
-
Paper title:Modeling and training strategies for language recognition systems
-
Paper track:4.1 Language identification and verification, lang/Oral Presentation
-
Paper status:Accept
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Raphaël Duroselle | 2008 NIST Speaker Recognition Evaluation | /N |
Documentation:
None
Speech
Corpus,
Language Type:
Monolingual
Languages:
Arabic Bengali Dari Egyptian Arabic English Georgian Hindi Iranian Persian Italian Japanese Khmer Korean Lao Mandarin Chinese Min Nan Chinese Moroccan Arabic Panjabi Persian Russian Spanish Tagalog Thai Tigrinya Urdu
Availability:
From Owner
License:
LDC
Size:
950 hoursProduction Status:
Existing-updated
Use:
Language Identification
-
Paper title:Modeling and training strategies for language recognition systems
-
Paper track:4.1 Language identification and verification, lang/Oral Presentation
-
Paper status:Accept
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Raphaël Duroselle | 2008 NIST Speaker Recognition Evaluation Training Set Part 2 | /N |
Documentation:
None
Speech
Corpus,
Language Type:
Multilingual
Languages:
Cantonese English French German Gishu Greek Gujarati Hebrew Hindi Indonesian Japanese Korean Mandarin Persian Portuguese Runyankore Russian Spanish Turkish Vietnamese
Availability:
Freely Available
License:
OpenSource
Size:
22.8 GByte Production Status:
Newly created-in progress
Use:
Speech Recognition/Understanding
-
Paper title:Speaking rate, information density, and information rate in first-language and second-language speech
-
Paper track:1.10 Bilingual and L2 acquisition and processing/Oral Presentation
-
Paper status:Accept - Poster
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Ann Bradlow | The ALLSSTAR Corpus | /N |
Documentation:
Documentation in English is available to the public (via the project website)
Written
Corpus,
Language Type:
Multilingual
Languages:
Bengali Hindi Kannada Sanskrit Telugu Urdu
Availability:
Freely Available
License:
Size:
None MByte Production Status:
Existing-used
Use:
-
Paper title:Analysing cross-lingual transfer in lemmatisation for Indian languages
-
Paper track:Short paper/
-
Paper status:Accept Poster
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Kumar Saunack | SIGMORPHON 2019 Shared Task 1 dataset | /N |
Documentation:
None
Written
Corpus,
Language Type:
Multilingual
Languages:
Hebrew Hindi Polish
Availability:
Freely Available
License:
Creative Commons
Size:
277.701 sentences Production Status:
Existing-used
Use:
Extraction of Multiword Expressions
-
Paper title:Verbal Multiword Expression Identification: Do We Need a Sledgehammer to Crack a Nut?
-
Paper track:Long paper/
-
Paper status:Accept Poster
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Carlos Ramisch | PARSEME corpora | /N |
Documentation:
None
Written
Lexicon,
Language Type:
Multilingual
Languages:
'Auhelawa Abau Aceh Achang Acholi Achuar-Shiwiar Aché Adamawa Fulfulde Adele Adhola Adi Adioukrou Aekyom Afrikaans Agarabi Aguacateco Aguaruna Agusan Manobo Agutaynen Aimol Ajië Ajyíninka Apurucayali Akawaio Akeu Akha Akoose Alamblak Alangan Alekano Algonquin Alladian Alune Alur Ama Amanab Amarakaeri Amarasi Ambai Ambulas Amele Amganad Ifugao Amharic Amri Ancient Greek Aneme Wake Angaatiha Angal Heneng Angami Naga Angor Anjam Anufo Ao Naga Apalaí Apatani Apinayé Apurinã Arabela Arifama-Miniafia Armenian Arop-Lukep Arosi Aruamu Asháninka Ashéninka Pajonal Assyrian Neo-Aramaic Ata Manobo Atatláhuca Mixtec Au Aukan Avar Avokaya Awa Awa-Cuaiquer Awadhi Awiyaana Ayoreo Ayutla Mixtec Azerbaijani Baatonum Baba Malay Bafia Bafut Bahasa Melayu Baka Bakairí Balangao Balantak Bali Bamanankan Bana Bandial Banggai Baoulé Barai Barasana Bargam Bariai Baruya Bashkir Basque Bassari Batad Ifugao Batak Angkola Batak Dairi Batak Karo Batak Simalungun Batak Toba Bauzi Bavarian Bawm Chin Bedjond Beembe Bekwarra Belarusan Belize Kriol English Bemba Bembe Benabena Bengali Berom Bete-Bendi Biangai Biatah Biete Bima Bimin Bimoba Bine Binukid Binumarien Bislama Bissa Bisu Bokmal Norwegian Boko Bokobaru Bola Bomu Bora Border Kuna Borong Bribri Buamu Bugawac Bugis Buhid Bukiyip Bulgarian Buli Bulu Bumbita Arapesh Bunama Burarra Burmese Burum-Mindik Busa Cabécar Cacua Caluyanun Cameroon Mambila Camsá Candoshi-Shapra Capanahua Caquinte Car Nicobarese Carapana Carib Caribbean Hindustani Caribbean Javanese Carrier Cashibo-Cacataibo Cashinahua Casiguran Dumagat Agta Catalan-Valencian-Balear Cavineña Cebuano Central Aymara Central Bicolano Central Bontok Central Cagayan Agta Central Cakchiquel Central Dusun Central Huasteca Nahuatl Central Khmer Central Kurdish Central Mnong Central Yupik Central-Eastern Niger Fulfulde Cerma Chachi Chamacoco Chamorro Chang Naga Chavacano Chayahuita Chayuco Mixtec Chechen Cherokee Chhattisgarhi Chipaya Chiquihuitlán Mazatec Chiquitano Chiripá Chopi Chortí Chothe Naga Chuave Chumburung Chuukese Chuvash Chácobo Cishingini Coatlán Mixe Coatzospan Mixtec Cofán Cogui Colorado Comaltepec Chinantec Coptic Cornish Cotabato Manobo Croatian Cubeo Cuiba Culina Czech Dadibi Daga Dan Dangaléat Danish Dano Dawawa Dawro Dedua Deg Denya Desano Dhao Dibabawon Manobo Digo Dii Dimasa Djambarrpuyngu Djimini Senoufo Dobu Dogrib Doyayo Dupaninan Agta Duri Duruma Dutch East Kewa Eastern Bolivian Guaraní Eastern Bontok Eastern Bru Eastern Canadian Inuktitut Eastern Highland Chatino Eastern Huasteca Nahuatl Eastern Jacalteco Eastern Kanjobal Eastern Krahn Eastern Mari Eastern Oromo Eastern Tawbuid Efik Ejagham Ekajuk El Nayar Cora Endo English Enxet Erzya Ese Ese Ejja Esperanto Estonian Ewage-Notu Ezaa Faiwol Falam Chin Farefare Faroese Fasu Fijian Fijian Hindustani Filipino Finnish Fon Fore French Ga Ga'dang Gagauz Galela Galo Adi Gamo Ganda Gangte Garifuna Garo Gbagyi Gen Georgian Gheg Albanian Ghomálá' Gidar Gikuyu Gikyode Girawa Gofa Gogo Gokana Golin Gonja Gor Gorontalo Gourmanchéma Greek Guahibo Guajajára Guambiano Guanano Guarayu Guayabero Gude Guerrero Amuzgo Guerrero Nahuatl Guhu-Samane Guinea Kpelle Gujarati Gulay Gumatj Gumuz Gun Gusii Gwahatike Gwich'in Haitian Creole French Haka Chin Hakka Chinese Halh Mongolian Halia Hamer-Banna Hanga Hanunoo Hausa Hawai'i Creole English Haya Hebrew Hehe Helong Highland Oaxaca Chontal Highland Puebla Nahuatl Hiligaynon Hindi Hiri Motu Hixkaryána Hmong Daw Hmong Njua Hopi Hote Hrangkhol Huambisa Huautla Mazatec Huichol Huli Hungarian Iatmul Iban Ibatan Icelandic Igbo Ignaciano Ika Ikwere Ikwo Ilianen Manobo Ilocano Imbongu Inabaknon Indonesian Inga Inoke-Yate Iraqw Iraya Iriga Bicolano Irigwe Irish Gaelic Islander Creole English Isnag Isthmus Mixe Isthmus-Mecayapan Nahuatl Italian Itawit Iu Mien Ivatan Ivbie North-Okpela-Arhe Iwal Iyo Iyo'wujwa Chorote Iyojwa'ja Chorote Izere Izii Jalapa de Díaz Mazatec Jamaican Creole English Jamiltepec Mixtec Japanese Jarai Javanese Jingpho Jola-Fonyi Jola-Kasa Jukun Takum Jula Juquila Mixe Jur Modo Kabiyé Kabyle Kadiwéu Kafa Kagayanen Kagulu Kahua Kaingáng Kaiwá Kako Kalagan Kalam Kalanga Kamano Kamasau Kambaata Kamwe Kandawo Kanite Kankanaey Kannada Kapingamarangi Kara Karachay-Balkar Karajá Karakalpak Karamojong Karbi Kasua Kayabí Kazakh Keapara Kein Kekchí Kele Keley-I Kallahan Kenga Kenyang Keyagana Khakas Khiamniungan Naga Kim Kimré Kinaray-A Kire Kirghiz Kiribati Kisar Kituba Kobon Kom Komba Komi-Zyrian Komso Konai Konni Kono Konyak Naga Koongo Koorete Korafe Korean Koreguaje Koronadal Blaan Kosena Kouya Koya Krio Kuanua Kube Kukele Kuku-Yalanji Kumam Kuman Kumyk Kunimaipa Kuot Kupang Malay Kupsabiny Kuranko Kusaal Kutep Kutu Kuwaa Kuwaataay Kwaio Kwanga Kwanyama Kwara'ae Kwere Kwoma Kyaka Laari Lacandon Ladakhi Lahu Lahu Shi Lalana Chinantec Lama Lamba Lambya Lamkang Lampung Lango Lao Lashi Latin Latvian Lealao Chinantec Ledo Kaili Lega-Mwenga Lelemi Lengua Lenje Lewo Lhomi Liangmai Naga Limbu Limbum Limos Kalinga Lingala Literary Chinese Lithuanian Lobi Loma Low Saxon Lozi Luang Lukpa Luo Luwo Lyélé Ma'anyan Ma'di Maasina Fulfulde Mabaan Maca Macedonian Machame Machiguenga Macuna Macushi Mada Madak Madura Mafa Maithili Maiwa Makaa Makasar Makonde Malayalam Malba Birifor Male Maltese Mamanwa Mamara Senoufo Mamasa Mampruli Manam Mandinka Mangga Buang Manggarai Mangseng Manikion Mankanya Mansaka Maori Mape Mapos Buang Mapudungun Maram Naga Maranao Marathi Marba Marik Maring Naga Marshallese Maru Masaba Masana Masbatenyo Maskelynes Matal Matigsalug Manobo Matsés Mauwake Maxakalí Mayo Mayoyao Ifugao Mazahua Central Mazatlán Mixe Mbay Mbo-Ung Mbuko Mbula Mbunda Mbyá Guaraní Mekeo Melpa Mende Mengen Mentawai Merey Meyah Mian Michoacán Nahuatl Micmac Middle English Min Nan Chinese Minangkabau Minaveha Minica Huitoto Mizo Moba Mocoví Mofu-Gudur Mokole Molima Mong Leng Mong Njua Mongo-Nkundu Mongondow Mono Moose Cree Mopán Maya Morisyen Moro Moskona Motu Mountain Koiali Moyon Naga Mufian Muinane Mumuye Muna Mundang Mundani Mundurukú Murle Murui Huitoto Musey Muyang Mískito Mòoré Mün Chin Mündü Naasioi Nabak Nadëb Nafaanra Nakanai Nalca Nama Nande Nandi Naro Navajo Nawdm Ndamba Ndau Ndebele Ndo Ndogo Ndonga Nepali Nga La Ngaju Ngangam Ngawn Chin Ngiemboon Ngindo Ngiti Ngombe Ngulu Ngäbere Nias Nigeria Mambila Nigerian Fulfulde Nii Nilamba Ninzo Nivaclé Nkonya Nobonob Nocte Naga Nogai Nomaande Nomatsiguenga Noone Nopala Chatino North Alaskan Inupiatun North Mofu Northeastern Dinka Northern Dagara Northern Emberá Northern Grebo Northern Khmer Northern Kissi Northern Kurdish Northern Mam Northern Oaxaca Nahuatl Northern Puebla Nahuatl Northwest Alaska Inupiatun Northwest Gbaya Ntcham Numanggang Nyindrou Nyishi Nynorsk Norwegian Obolo Ocotepec Mixtec Ogea Old Church Slavonic Olusamia Ozumacín Chinantec Palantla Chinantec Pamplona Atta Paraguayan Guarani Patpatar Pele-Ata Peñoles Mixtec Phom Naga Pichis Ashéninka Pinotepa Nacional Mixtec Plapo Krumen Psikye Pular Qaqet Quiotepec Chinantec Rabinal Achí Russia Buriat Rwanda S'gaw Karen Sa'a Saamia Sabu Safeyoka Saint Lucian Creole French Samba Leko San Blas Kuna San Jerónimo Tecóatl Mazatec San Juan Colorado Mixtec San Juan Cotzal Ixil San Mateo del Mar Huave San Miguel el Grande Mixtec San Pedro Amuzgos Amuzgo San Sebastián Coatán Chuj Santa María Zacatepec Mixtec Santa Teresa Cora Sar Sarangani Blaan Sarangani Manobo Sateré-Mawé Sea Island Creole English Sekpele Sepik Iwam Seselwa Creole French Sharanahua Shuar Silacayoapan Mixtec Siyin Chin Sochiapan Chinantec South Fali South Giziga Southern Altai Southern Birifor Southern Bobo Madaré Southern Carrier Southern Ghale Southern Kalinga Southern Kisi Southern Nambikuára Southern Nuni Southern Puebla Mixtec Southwest Gbaya Southwestern Dinka Standard Arabic Standard German Tabasco Chontal Tabo Tagabawa Takuu Tangkhul Naga Tataltepec Chatino Tedim Chin Tenango Nahuatl Tepetotutla Chinantec Tepeuxila Cuicatec Tetelcingo Nahuatl Teutila Cuicatec Tezoatlán Mixtec Tlahuitoltepec Mixe Tol Toro So Dogon Totontepec Mixe Toura Tsikimba Tsimané Tuma-Irumu Tumbalá Chol Tungag Tuwali Ifugao Uab Meto Ucayali-Yurúa Ashéninka Umanakaina Umiray Dumaget Agta Una Usila Chinantec Vengo Veracruz Huastec Waimaha Wancho Naga Wandala Waorani Wayuu Welsh West Kewa West-Central Limba Western Apache Western Arrarnta Western Bolivian Guaraní Western Bukidnon Manobo Western Frisian Western Highland Chatino Western Huasteca Nahuatl Western Kanjobal Western Niger Fulfulde Wichí Lhamtés Güisnay Wichí Lhamtés Nocten Wipi Woun Meu Xaasongaxango Yabem Yanesha' Yocoboué Dida Yosondúa Mixtec Zaiwa Zarma Zemba Zotung Chin Zulgo-Gemzek Éwé Ömie
Availability:
Freely Available
License:
CC BY-NC-ND license (Attribution-NonCommercial-NoDerivs)
Size:
7 MByte Production Status:
Newly created-finished
Use:
Opinion Mining/Sentiment Analysis
-
Paper title:UniSent: Universal Adaptable Sentiment Lexica for 1000+ Languages
-
Paper track:Terminology/poster presentation
-
Paper status:Accept Poster
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Main Contact | Ehsaneddin Asgari | UniSent | /N |
Documentation:
None
Written
Corpus,
Language Type:
Multilingual
Languages:
English Hindi Punjabi Tamil
Availability:
From Data Center(s)
License:
TDIL, Government of India
Size:
600000 sentences Production Status:
Existing-used
Use:
Machine Translation, SpeechToSpeech Translation
-
Paper title:Issues in chunking parallel corpora: mapping Hindi-English verb group in ILCI
-
Paper track:Short Paper
-
Paper status:Accept
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Author 1 | Esha Banerjee | Jawaharlal Nehru University | IN |
| Author 2 | Akanksha Bansal | Jawaharlal Nehru University | None |
| Author 3 | Girish Jha | Jawaharlal Nehru University, New Delhi | IN |
| Main Contact | Esha Banerjee | Jawaharlal Nehru University | None |
Documentation:
<Not Specified>
Written
Corpus,
Language Type:
Multilingual
Languages:
Bengali Hindi Tamil Telugu
Availability:
Freely Available
License:
OpenSource
Size:
2 MByte Production Status:
Newly created-finished
Use:
Machine Translation, SpeechToSpeech Translation
-
Paper title:No more beating about the bush : A Step towards Idiom Handling for Indian Language NLP
-
Paper track:Written
-
Paper status:Accept Oral
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Author 1 | Ruchit Agrawal | FBK, Trento | IT |
| Author 2 | Vighnesh Chenthil Kumar | IIIT Hyderabad | IN |
| Author 3 | Vigneshwaran Muralidaran | International Institute of Information Technology Hyderabad | IN |
| Author 4 | Dipti Sharma | IIIT, Hyderabad | IN |
| Main Contact | Ruchit Agrawal | FBK, Trento | None |
Documentation:
Available in EnglishLanguage Type:
Multilingual
Languages:
Belarusian English Hebrew Hindi
Availability:
Freely Available
License:
<Not Specified>
Size:
14 KByte Production Status:
Newly created-in progress
Use:
As annotation guidelines
-
Paper title:A Unified Annotation Scheme for the Semantic/Pragmatic Components of Definiteness
-
Paper track:Written
-
Paper status:Accept Oral
| Author Number | Name | Affiliation | Country |
|---|---|---|---|
| Author 1 | Archna Bhatia | Carnegie Mellon University | US |
| Author 2 | Mandy Simons | Carnegie Mellon University | US |
| Author 3 | Lori Levin | Carnegie Mellon University | US |
| Author 4 | Yulia Tsvetkov | Carnegie Mellon University | US |
| Author 5 | Chris Dyer | Carnegie Mellon University | US |
| Author 6 | Jordan Bender | University of Pittsburgh | US |
| Main Contact | Archna Bhatia | Florida Institute for Human and Machine Cognition | None |
Documentation:
<Not Specified>




